Papers with language family

8 papers
Massively Multilingual Adversarial Speech Recognition (N19-1)

Copied to clipboard

Challenge: Prior work in multilingual and cross-lingual speech recognition has been limited to a subset of the world's most-spoken languages.
Approach: They propose to use phonemes and phonemes as pretraining objectives to encourage language-independent representations.
Outcome: The proposed model is able to learn language-independent representations of speech using multilingual training.
The Geometry of Multilingual Language Model Representations (2022.emnlp-main)

Copied to clipboard

Challenge: XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning.
Approach: They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language.
Outcome: The proposed model can extract features for downstream tasks and cross-lingual transfer learning.
Discovering Representation Sprachbund For Multilingual Pre-Training (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models perform poorly on many languages and cross-lingual tasks due to typological differences and contradictions between some languages.
Approach: They propose to pre-train multilingual pre-trained models to handle cross-lingual tasks in one model.
Outcome: The proposed model improves performance on cross-lingual tasks compared to baselines on multiple languages .
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)

Copied to clipboard

Challenge: We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script.
Approach: They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic.
Outcome: The proposed training recipe improves multilingual transliteration for Indic languages.
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to detect cross-linguistic associations are not effective, but their effects are minor.
Approach: They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon.
Outcome: The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average).
Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource Scenarios (2020.emnlp-main)

Copied to clipboard

Challenge: a limited set of translations into one or more high-resource languages are available for POS tagging . a bi-LSTM architecture that uses contextualized word embeddings improves performance .
Approach: They propose an unsupervised cross-lingual transfer approach for part-of-speech tagging . they use the Bible as parallel data to learn POS taggers for target languages .
Outcome: The proposed approach improves accuracy on 12 diverse languages . the Bible is used as a parallel corpus for the study .
Towards a Better Understanding of Variations in Zero-Shot Neural Machine Translation Performance (2023.emnlp-main)

Copied to clipboard

Challenge: Prior work has investigated causes of poor zero-shot performance, but new study suggests it does not exhibit poor zero shot capability.
Approach: They propose to investigate the presence of significant variations in zero-shot performance . target-side translation quality is most influential factor, with vocabulary overlap impacting zero- shot capabilities .
Outcome: The results show that the target side translation quality is the most influential factor . linguistic properties, such as language family and writing system, play a role .
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are restricted to high- or mid-resource languages, and evaluate performance on higher-order tasks in reasoning and generation.
Approach: They propose a multilingual benchmarking tool to evaluate lexical comprehension and generation abilities of large language models.
Outcome: The proposed benchmarks cover 2700+ languages and surpasses existing benchmarks in terms of language coverage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations